Big Data Startups funded by Y Combinator (YC) 2026

September 2026

Browse 28 of the top Big Data startups funded by Y Combinator.

We also have a Startup Directory where you can search through over 5,000 companies.

  • Lyon
    Lyon
    Y Combinator LogoF2026
    Active • 2 employees • San Francisco
    Lyon builds private foundation models for banks, insurance companies, and fintechs. Each model learns from the company’s proprietary transactions, payments, clicks, and customer interactions to predict credit, fraud, collections, income, churn, and lifetime value; and runs inside the company’s cloud, so its data and intelligence remain under its control. Lyon is already working with a major insurer and trained a model on 28B transactions for a fintech serving tens of millions of active users, identifying premium-card converters 4x more precisely than its existing rules; the model is now being deployed for credit.
    machine-learning
    data-science
    finance
    big-data
  • Atlia
    Atlia
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    We use AI agents to manage short-term rental properties by running operations completely autonomously. We cut traditional human management fees in half so owners take home 20% to 30% more revenue. And we operate with near-zero boots-on-the-ground so we can scale nationwide with software margins and at venture speed in a traditionally physical industry.
    real-estate
    big-data
    artificial-intelligence
    ai
    proptech
  • Maingen
    Maingen
    Y Combinator LogoS2026
    Active • 2 employees • San Francisco
    Maingen builds simulations for industrial operations ($5 trillion slice of the US economy) so frontier labs can train models that actually run power plants, factories, and datacenters.
    artificial-intelligence
    ai
    b2b
    big-data
    energy
  • Hub
    Hub
    Y Combinator LogoP2026
    Active • 10 employees • San Francisco
    Hub is a multimodal lab providing real world data to frontier labs and robotic companies.
    ai
    big-data
    infrastructure
    crowdsourcing
    robotics
  • Parse
    Parse
    Y Combinator LogoF2025
    Active • 4 employees • San Francisco
    Parse lets developers build APIs for any website. Unlike traditional scraping tools that rely on slow, expensive headless browsers, Parse examines websites' underlying network requests and generates API endpoints automatically. This makes Parse 10-100x faster and cheaper than browser automation,.
    developer-tools
    api
    big-data
  • Sciloop
    Sciloop
    Y Combinator LogoF2025
    Active • 6 employees • San Francisco
    Sciloop creates expert-level math and physics problems that frontier AI models can't solve, then sells the data to AI labs for training and evaluation. Our problems are created by IPhO and IMO medalists — the top 0.01% of STEM talent globally. On our benchmark, models like GPT 5.4 Pro and Gemini 3.1 Pro score 0-5% on our hardest problems. We work with AI labs to supply continuous, fresh training data that pushes the frontier of mathematical and scientific reasoning. Founded by Bilal and Osman, International Physics Olympiad medalists from MIT with hands-on ML research experience at MIT CSAIL.
    ai
    big-data
    data-labeling
    marketplace
  • Zavo
    Zavo
    Y Combinator LogoF2025
    Active • 8 employees • London
    Zavo is a data company building proprietary, real-world data infrastructure for frontier AI. We partner with hospitals, businesses, operators and expert networks globally to access hard-to-reach data that cannot be found on the public internet. We source, structure, annotate and license this data for model training, post-training and evaluation, and use it to build specialised benchmarks and reinforcement learning environments grounded in real-world workflows, expert decisions and outcomes. Our focus spans healthcare and robotics, where high-quality proprietary data, realistic environments and reliable evaluations are critical to building increasingly capable AI systems.
    b2b
    ai
    big-data
  • Pleom
    Pleom
    Y Combinator LogoS2025
    Active0 • New York City
    Automatic insights on all your company data. Speed up business workflows to visualizations with a 0 learning curve.
    generative-ai
    artificial-intelligence
    ai
    analytics
    big-data
  • Liva AI
    Liva AI
    Y Combinator LogoS2025
    Active • 6 employees • San Francisco
    Helping to build more socially intelligent AI, starting with voice.
    b2b
    big-data
    data-labeling
    marketplace
    artificial-intelligence
  • ParaQuery
    ParaQuery
    Y Combinator LogoP2025
    Active • 1 employees • New York City
    ParaQuery is a fully-managed GPU-accelerated Spark solution providing double the performance at half the cost compared to solutions like Databricks and BigQuery. It is fully Spark-compatible, cloud agnostic, and has seamless integrations everywhere, all without vendor lock-in. ParaQuery makes it trivial to run large SQL and Spark workloads efficiently and serverlessly, without any data migrations.
    analytics
    infrastructure
    developer-tools
    big-data
    enterprise-software
  • AfterQuery
    AfterQuery
    Y Combinator LogoW2025
    Active • 30 employees • San Francisco
    AfterQuery is an applied research lab curating data solutions for frontier foundation model development. Serving every frontier AI lab.
    b2b
    artificial-intelligence
    ai
    big-data
    data-labeling
  • Vortexify
    Vortexify
    Y Combinator LogoF2024
    Active • 2 employees • New York City
    Vortexify is an AI app builder for full-stack, internal business apps. Businesses such as manufacturers, hospitals and logistics services use Vortexify to solve complex, operational problems such as scheduling, process optimization, forecasting, root cause analysis and scenario planning. Our agents encode business logic by mapping operational problems to the data in your systems and collaborating with domain experts. Operations teams can build apps without writing code. Engineering teams can edit and manage agent written code at any time. Vortexify builds applications beyond simple dashboards with native support for databases, scheduled tasks, optimization solvers, machine learning models, data pipelines, API calls and more. Our secure, managed infrastructure enables our agents to build quickly and get it right the first time. Solve your hardest operations problems today! Create a free account to get started. For additional technical services or support, contact us or book a demo.
    ai-assistant
    artificial-intelligence
    supply-chain
    operations
    big-data
  • Archil
    Archil
    Y Combinator LogoF2024
    Active • 11 employees • San Francisco
    Archil transforms S3 buckets into a 30x faster, unlimited, local disk. Archil enables AI, analytics, and serverless applications to instantly access massive data sets without waiting for data transfer. Researchers use Archil for shareable, local storage of data set and model versions that never runs out of capacity.
    ai
    developer-tools
    infrastructure
    machine-learning
    big-data
  • Voker
    Voker
    Y Combinator LogoS2024
    Active • 6 employees • Los Angeles
    Voker is the Agent Analytics Platform for monitoring and improving your AI agents. Companies like Dutch.com use our SDK to build better agents. Alex and Tyler met at a high-growth E-Commerce startup where Tyler was running Technology & Data. Together, they built AI products that bootstrapped the company profitably to $100MM in Revenue.
    analytics
    big-data
    developer-tools
    ai
    generative-ai
  • Sharpe
    Sharpe
    Y Combinator LogoS2024
    Active • 3 employees • San Francisco
    Sharpe helps traders go from idea to profit in minutes with AI, bundling petabytes of market data with high-performance infrastructure.
    ai
    finance
    analytics
    data-engineering
    big-data
  • Upsolve AI
    Upsolve AI
    Y Combinator LogoW2024
    Active • 5 employees
    Upsolve AI is the analytics platform for deploying governed, grounded, and trustworthy data agents that know your business. Data teams use Upsolve's Agent Studio to get to AI-ready data foundations fast, unify and curate their context layer, and test, tune and monitor their Upsolve Data Agent to their benchmark. The company is founded by Ka Ling Wu and Serguei Balanovich, who built a similar product at Palantir before (featured in Palantir's S-1), growing it to 50+ enterprise customers and 8-figures of annual revenue in 2 years.
    analytics
    developer-tools
    big-data
    b2b
    ai
  • Energent AI
    Energent AI
    Y Combinator LogoS2023
    Active • 10 employees • San Francisco
    Every AI tool will confidently hand you a wrong number — and in engineering and finance, a wrong number is expensive. Energent.ai is an AI teammate that does your data work on a secure desktop: reading documents and engineering drawings, cleaning messy data, writing code, building charts. Then a second, independent agent audits that output against your original source files and flags every claim that doesn't hold up — catching roughly 3x more hallucinations than the model alone in our internal evals. Engineering teams check drawing sets before issue. Finance teams tie numbers back to source. The best teams encode their own checking standards as reusable rules, so it checks the way their firm checks.
    ai-assistant
    automation
    big-data
    productivity
    enterprise-software
  • NewsCatcher
    NewsCatcher
    Y Combinator LogoS2022
    Active • 22 employees • Kyiv, Ukraine, 02000
    CatchAll by NewsCatcher is a recall-first web search API built for queries where the results are spread across hundreds or thousands of pages on the web. Instead of returning the top ranked links like traditional search engines, CatchAll retrieves a large candidate set from the web, validates which pages actually match the query, and extracts structured records of real-world events. Developers and data teams use CatchAll to answer “long-list” questions such as tracking regulatory actions, funding rounds, product launches, corporate expansions, or cybersecurity incidents. The output is not just links but clean, deduplicated datasets that can power AI agents, monitoring systems, analytics pipelines, and market intelligence workflows. CatchAll runs on the data infrastructure developed by NewsCatcher, which continuously indexes millions of articles and public web pages across a global network of sources.
    big-data
    enterprise
    enterprise-software
    saas
    artificial-intelligence
  • Endla
    Endla
    Y Combinator LogoS2021
    Active • 8 employees • Brisbane QLD, Australia
    Endla increases the value of oil & gas wells by providing software that helps design the well that maximizes ROI. Our product AlphaSpace, optimizes the well design by automatically producing many high-quality options which the engineer can then measure against their business objectives (auto-design). Having software that finds the optimal solution, empowers the engineer to work a layer up on understanding the problem and specifying the important objectives. Our vision is to make auto-design and auto-operation software part of the workflow for every engineer working with physical assets.
    energy
    big-data
    analytics
  • Supabase
    Supabase
    Y Combinator LogoS2020
    Active • 120 employees • San Francisco
    Supabase is the easiest way to get started with Postgres. Each project within Supabase is an isolated Postgres cluster, allowing customers to scale independently, while still providing the features that you need to build: instant database setup, auth, row level security, realtime data streams, auto-generating APIs, and a simple to use web interface. We are 100% remote.
    open-source
    databases
    data-engineering
    big-data
    developer-tools
  • Gecko Robotics
    Gecko Robotics
    Y Combinator LogoW2016
    Active • 230 employees • Pittsburgh, PA, USA
    Gecko Robotics is the pioneer of AI + Robotics [AIR technology], transforming how the world builds, operates, and maintains its most critical infrastructure for a more reliable and sustainable future. Using fixed sensors and robots that climb, crawl, swim, and fly, we combine first-order data layers with the predictive power of AI into a single source of truth for the physical world. Cantilever™ is our operating platform, powered by AIR technology, that empowers teams to achieve operational excellence through actionable data for immediate and long-term planning.
    big-data
    energy
    robotics
    data-engineering
    artificial-intelligence
  • Deasy Labs
    Deasy Labs
    Y Combinator LogoS2023
    Acquired • 8 employees • New York City
    Deasy Labs was acquired by Collibra in July 2025 (global leader in enterprise data governance). Deasy Labs provides metadata orchestration for AI workflows. Deasie's platform provides the best way for AI teams to create and embed high-quality, customized metadata into their AI workflows (e.g., RAG, Agentic frameworks). Our three founders (from Amazon, McKinsey/QuantumBlack & MIT) previously built an ML data governance tool from 0 to 1 within McKinsey, which we deployed with 11 Fortune 500 companies. We saw in early 2023 the ability to create high-quality metadata (without reliance on domain experts) would be a key factor in achieving the accuracy & speed in GenAI applications required for production. Our investors include General Catalyst, Y Combinator, RTP Global and world experts in enterprise data. Website: https://deasylabs.com
    ai-assistant
    data-labeling
    databases
    big-data
    artificial-intelligence
  • Tarsal
    Tarsal
    Y Combinator LogoS2021
    Acquired • 10 employees • New York City
    Tarsal is a data pipeline custom built for security teams. As security data grows 25% year over year, security teams desperately need access to best-in-class data infrastructure. Tarsal bridges the gap between the modern data stack and security teams, pioneering the modern security data stack.
    cybersecurity
    data-engineering
    big-data
    b2b
  • Terark
    Y Combinator LogoW2017
    Acquired • 2 employees • Beijing, China
    Terark built a new storage engine for Database and Data Systems. Our technology enables direct search on highly compressed data, with 200X faster read performance and more than 10X storage savings (better than Google's LevelDB, Facebook's RocksDB), getting larger scalability with lower cost for big data applications. Alibaba is our paying customer, and we are a YCombinator company.
    big-data
    cloud-computing
  • Scuba
    Scuba
    Y Combinator LogoW2013
    Acquired • 51 employees
    Scuba is the fast and scalable event-based analytics solution to answer critical business questions about how customers behave and products are used. Interana allows users to analyze and explore the key business metrics that matter most in a data-driven world – such as growth, retention, conversion and engagement – in seconds, rather than the hours or days it often takes with existing solutions. Interana allows customers to discover and investigate these key insights easily through its visual and interactive interface, which makes data analysis a natural extension of everyone’s workflow.
    analytics
    big-data
    data-engineering
    data-visualization
  • Mattermark
    Y Combinator LogoS2012
    Acquired • 11 employees • San Francisco
    At Mattermark, we’re accelerating sales and deal making through data and automation. Mattermark collects and organizes comprehensive information on the world’s fastest growing companies. In minutes, get actionable data that lets you pinpoint the companies and people you need to know or do business with. Today, over 500 companies use Mattermark to discover high quality leads, prioritize prospects and increase conversion rates.
    big-data
    investing
  • Citus Data
    Citus Data
    Y Combinator LogoS2011
    Acquired • 45 employees • San Francisco
    The amount of time businesses spend on their databases is altogether too much time. Citus is fixing this problem. Citus is worry-free Postgres. Built to scale out, Citus is an extension to Postgres that is available as open source, as enterprise software that can be run on-prem or on any cloud, and as a fully-managed database as a service. Whether you have a multi-tenant application that needs to scale out, or you need performance for your real-time analytics customers, with Citus, you can focus on your app—not your database. Founded in 2011, Citus Data is venture backed by Khosla Ventures and Data Collective. Citus is a Y Combinator alumnus and has offices in San Francisco’s SoMa district and Istanbul, Turkey. At Citus, we make it simple to scale out Postgres. Citus Data online: www.citusdata.com Documentation: docs.citusdata.com GitHub: github.com/citusdata/citus
    databases
    big-data
    open-source
  • Amiato
    Y Combinator LogoW2012
    Acquired • 2 employees • San Francisco
    Amiato's real-time integration service moves your data to where it's most valuable to you. Today's flexible databases like MongoDB and CouchBase let agile businesses accelerate and scale their operations, but analyzing their data has been elusive. Amiato unlocks the value of that unstructured data by bridging the gap to familiar tools in the rich structured world of BI, all with zero setup work. Run reports, do interactive ad-hoc analysis, and combine disparate silos. Our Schema-lift (TM) technology allows you to immediately integrate new endpoints and seamlessly keep up with changes in your data. We get customers up and running in a day instead of weeks, and let them focus on business instead of wrestling infrastructure.
    big-data
    analytics